Systematic Biology
◐ Oxford University Press (OUP)
Preprints posted in the last 7 days, ranked by how well they match Systematic Biology's content profile, based on 144 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.
Zhou, Y.; Gong, L.; Niu, G.; Shi, H.; Gutell, R.; Li, X.; Wei, M.
Show abstract
Animal mitochondrial rRNAs are commonly viewed as structurally reduced, yet sponge mt SSU rRNAs range from compact to highly expanded structures. Using nine conserved structural anchors, we compared 216 taxonomically resolved records from four classes and 22 orders, including 16 freshwater Spongillida and 200 marine sponges. Twelve homologous hypervariable substructures were coded as structural types, and their ordered combinations as composite types. We identified 38 structural types and 62 composite types across molecules ranging from 853 to 2,019 nt. Hexactinellida and freshwater Spongillida were each uniform for a distinct composite type but differed markedly in overall structure: hexactinellid mt SSU rRNAs were compact, whereas those of Spongillida were long and contained four to five candidate insertion regions. These results show that a conserved scaffold can accommodate extensive lineage-associated structural variation and provide a practical framework for comparing highly divergent mitochondrial rRNAs.
Zeng, Z.; Wang, Y.
Show abstract
Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [≥]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).
Bohnenkaemper, L.; Stoye, J.
Show abstract
The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.
Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.
Show abstract
Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.
Siemers, M.; Lopez, J. L.; Dutilh, B. E.
Show abstract
Bacteriophages can only be understood through their interactions with bacterial hosts. As environmental sequencing efforts expanded, the number of available phage genome sequences has exploded, yet the vast majority of these sequences lack host information. Predicting the host of a newly observed phage is therefore a key challenge in virology. Several computational tools can predict phage-host relationships from genomic data, but they share notable limitations: (1) the number of different hosts that can be predicted remains relatively restricted; (2) tools tend to assign confident host predictions to non-viral input sequences; and (3) most tools have a trade-off between accuracy and speed. Here we present PhageTransformer (PT), a deep learning model for phage-host prediction that addresses these limitations. We benchmark PT against existing tools on 3,881 independent phage-host pairs from GenBank and public HiC data, and demonstrate that it achieves competitive or superior prediction accuracy at greatly reduced runtime.
Alve, S. R.; Rahman, S.; Meem, S. M. A. C.
Show abstract
A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.
Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.
Show abstract
Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.
Duffin, P. J.; Ruggeri, M.; Conn, T.; Baums, I. B.; Blanco-Pimentel, M.; Bosch, P.; Carne, L.; Danser, N.; Montoya-Maya, P.; Morikawa, M.; Muller, E. M.; Winters, R. S.; Baker, A. C.; Cunning, R.; Dahlgren, C.; Parkinson, J. E.; Kenkel, C. D.
Show abstract
Genomic signatures can provide key insight into the evolutionary history and remaining adaptive potential of threatened populations. As demographic decline erodes both diversity and the processes maintaining it, understanding how remaining variation is distributed becomes increasingly important for conserving species like the staghorn coral, Acropora cervicornis, a foundational but critically endangered Caribbean reef-builder. We analyzed 46 high-coverage A. cervicornis genomes from 10 locations across the tropical western Atlantic to evaluate neutral and adaptive structure, genomic diversity, demographic history, inbreeding, and connectivity, and generated a regional haplotype reference panel for future genomic monitoring. Genome-wide analyses recovered recurring regional substructure, but differentiation was modest and partly explained by isolation-by-distance and spatial variation in effective migration. Subpopulations had similar levels of genomic diversity, shared demographic history, and limited evidence of local adaptation. These patterns support interpreting sampled Caribbean populations as a single evolutionarily significant unit (ESU) containing multiple regional management units (MUs), rather than as deeply divergent evolutionary lineages. Despite substantial retained variation and low current inbreeding, estimated contemporary effective population size was small, suggesting an increased vulnerability to the effects of drift as demographic collapse continues, especially if structure is reinforced by isolated management. Together, our findings emphasize the urgent need for interventions that preserve and enhance genomic diversity, including risk-managed assisted gene flow. Supported by the haplotype reference panel developed here, these strategies will require coordinated efforts across regional entities to conserve and restore A. cervicornis as a jointly managed, single ESU.
Baruah, G.; KC, Y. K.
Show abstract
The shape of density-dependence governs species persistence, and ecosystem stability. Yet, whether per-capita growth declines sublinearily, or superlinearily with density remains hotly debated. Growth rates across the tree of life have been shown to decline sublinearly with density, whereas theory founded on resource competition predicts the opposite. Here, we resolve this discrepancy and show that sublinearity can readily emerge from geometric constraints on consumer interactions. By linking inter individual spacing, movement and interference rates, we derive two limiting-interference regimes, one of which the well-mixed limit recovers the form of classic Beddington DeAngelis interference response. We then developed an individual-based model from first principles which reproduces the derived sublinearity response, and further use empirical data from published consumer-resource experiments that also bears the signature of sublinear density-dependence. Further, embedding the interference mechanisms underlying the emergence of sublinear density-dependence in coexistence theory opens a new regime for species coexistence where classical theory fails to predict. Our framework indicates that non-consumptive interactions are not merely a correction to resource competition but might be a distinct axis along which diverse communities may potentially coexist.
Dos Santos, M.; Ohtsuki, H.; Mullon, C.
Show abstract
Reputation plays a major role in supporting cooperation among unrelated individuals through indirect reciprocity. By helping others, individuals build a good personal reputation and receive greater benefits from future partners. Most models of indirect reciprocity assume that a person's reputation reflects only their own behaviour. Yet in many societies, people are also judged by their family's reputation. How family reputation affects the evolution of cooperation, and whether reliance on it can itself evolve, remain unclear. Here we show that reputation inheritance expands the conditions under which indirect reciprocity favours cooperation, increasing helping and favouring greater reciprocity. Greater reciprocity in turn favours stronger reliance on inherited reputation, creating a positive feedback that stabilises cooperation, especially when interactions are infrequent or personal behaviour is difficult to observe. This feedback arises because cooperation generates future benefits both for the individual, through their personal reputation, and for their descendants, through inherited reputation. Reputation inheritance thereby provides a route via which kin selection and reciprocity, often treated as alternative explanations for cooperation, can reinforce one another. Our model helps explain why family-based reputation occurs across diverse human societies and provides an evolutionary framework for studying phenomena organised around family standing, including kin-based institutions, feuds between families and honour-based violence within them.
Adam, L.; Montagna, M.; Roma, V.; Mancini, A.; Papafitsoros, K.
Show abstract
Wildlife re-identification (re-ID) is a widely used and powerful tool with diverse applications in animal ecology and conservation. Current automated methods typically operate on single images of a single body part of the animal. However, a single encounter may contain multiple images capturing different body regions, each providing complementary individual-specific information. In contrast to automated approaches, researchers often manually select the most suitable images and regions for identification based on factors like visibility, occlusion and image quality. This creates a mismatch between automated methods and field practice, limiting the practical adoption of current automated re-ID pipelines. Here, we address this by introducing an encounter-based, multi-body-part re-ID framework, using sea turtles as a model taxon. Our framework combines three elements: (1) An orientation-aware deep learning model, TurtleDetector, that in addition to the full bodies, it also automatically segments key body regions, i.e. heads, front and hind flippers, from images within an encounter; (2) a hybrid body-part-specific retrieval method, that sequentially combines a fast global-feature model (MiewID or DINOv3) with a more accurate but costlier local-feature model (ALIKED with LightGlue); and (3) a merged identity-prediction strategy that selects the highest calibrated similarity score across all available body parts and images of an encounter. We evaluate the framework on three long-term re-ID datasets spanning three species, loggerheads, greens, and hawksbill turtles, under an evaluation protocol that mirrors real-world, time-aware re-ID workflows. Across datasets, combining multiple body regions consistently improved identification performance over the best-performing single body region, resulting to an increase of 4-6% in top-1 accuracy. Interestingly, body regions traditionally underused in sea turtle re-ID, such as the hind flippers and carapaces, provided complementary identifying information that improved encounter-level re-ID when integrated through the hybrid retrieval method. Our findings demonstrate that automated wildlife re-ID can benefit from moving beyond single-image, single-body-part identification towards encounter-level integration of all available visual evidence. Our work further suggests that, where feasible, field photo-acquisition protocols should aim to capture multiple informative views of an individual during each encounter. Importantly, many species and taxa, including elephants, primates, cetaceans, and other large vertebrates, possess such individual-specific features across multiple body regions, highlighting the broad potential applicability of our framework.
Koshute, P.; Fagan, W. F.
Show abstract
Ecologists remotely track movement steps of animals (e.g., via global positioning systems) and use step selection functions to study the effect of environmental factors upon their movement decisions. Constructing such functions requires pairing each observed step with some number of unobserved but feasible comparison steps. Larger numbers of comparison steps generally yield better estimates but also incur potentially challenging computational demands. Thus, it is important to determine an appropriate number of comparison steps. No established guidance exists for this decision. Here, we use simulated tracks to assess how many comparison steps are needed, fitting each set of steps to a conditional logistic regression model. We monitor errors in estimated effects for several classes of tracks, identifying the number of comparison steps for which mean relative absolute error in estimated effects is consistently low. By this criterion, 32 comparison steps per observed step are needed for our primary class of simulated tracks. Tracks in more homogeneous landscapes, tracks with shorter mean step lengths, or shorter tracks generally require more comparison steps (ranging from 64 to 128 per observed step) to achieve the same level of accuracy. Longer tracks generally require fewer comparison steps (16 per observed step). These results clearly demonstrate that the number of comparison steps influences how well step selection functions estimate covariate effects and provides initial direction in a research area that currently lacks quantitative guidance. Movement ecologists should take care when selecting the number of comparison steps paired with each observed step because those decisions matter.
Laigaard, J.; Moeller, M. O.; Olsen, M. H.; Overgaard, S.; Mathiesen, O.; Karlsen, A. P. H.
Show abstract
Background: In Denmark, perioperative high-dose glucocorticoid treatment were step-wisely implemented for total hip arthroplasty (THA), total knee arthroplasty (TKA), and unicompartmental knee arthroplasty (UKA). We aimed to estimate the effect of a single high dose of glucocorticoids on opioid consumption following primary THA, TKA, and UKA. Methods: This was a prespecified analysis of a multicenter natural experiment using electronic health record data. We included elective THA, TKA, or UKA surgeries performed in Eastern Denmark from 2017-2025. At each center, surgeries before implementation of high-dose glucocorticoids served as controls, whereas surgeries after implementation comprised the intervention group. The primary outcome was the between-group difference in cumulative 0-24h opioid consumption, which included preemptive end-of-surgery doses. The predefined minimal important difference was set at 5 mg IV morphine equivalents. Secondary outcomes were maximum 0-10 numerical rating scale (NRS) pain score and incidence of opioid-related adverse events within 24 hours, hospital length of stay, and days alive and out of hospital at 30 days. Results: A total of 47,317 surgeries performed at nine centers were analyzed: 13,010 controls and 34,307 in the intervention group. During the study period, five centers implemented high-dose glucocorticoids for THA patients, two for TKA/UKA patients. High-dose glucocorticoids were administered to 6% of patients before implementation versus 92% after. High-dose glucocorticoids resulted in a mean reduction of 3.8 mg intravenous (IV) morphine equivalents (95% CI 3.3;4.3). The intervention also reduced the maximum 0-24h NRS pain score by 0.8 points (99% CI 0.7;0.9), but there was no difference in adverse events, length of stay, or days alive and out of hospital. Conclusions: Implementation of high-dose glucocorticoids reduced 0-24-hour opioid consumption by 3.8 mg IV morphine equivalents after elective hip and knee arthroplasty. This difference was below the prespecified minimal important difference threshold. Online registration: https://doi.org/10.1101/2025.11.11.25339982
Rohd, S. B.; Thorup, A. A.; Wilms, M.; Schiavon, M.; Streyma, D. H. B.; Laursen, A. F.; Bundgaard, A. F.; Sondergaard, A.; Krantz, M. F.; Veddum, L.; Hjorthoj, C.; Greve, A.; Mors, O.; Nordentoft, M.; Hemager, N.; Gregersen, M.
Show abstract
Objective: This study examined the prevalence of psychotic experiences (PE) and how early onset and persistence of PE contribute to risk and severity of mental disorders in adolescents at familial high-risk of schizophrenia (FHR-SZ) or bipolar disorder (FHR-BP) and adolescents from a population-based control group (PBC). Methods: This is the second follow-up of a nationwide cohort study including 522 children at FHR-SZ (N=202), FHR-BP (N=120), and PBC (N=200). Participants were assessed at ages 7, 11, and 15 using a semi-structured interview to evaluate PE and mental disorders. Results: At age 15, adolescents at FHR-SZ reported more PE than PBC over the past six months (current) and the past four years, while adolescents at FHR-BP only reported more current PE. PE reported at two or three timepoints (persistent PE) predicted any Axis I disorder in mid-adolescence, corresponding to three- (OR 2.9, 95% CI [1.5-5.7]) and 21-fold (OR 21.4, 95% CI [2.8-162.3]) increased risks, respectively. Persistent PE also predicted multimorbidity, with three- (OR 2.8, 95% CI [1.0-7.6]) and four-fold (OR 4.1, 95% CI [1.2-14.1]) increased risks, respectively. This was after adjustment for sex, early mental disorders, and familial risk. Conclusions: This study demonstrates a strong link between persistent PE and mid-adolescence mental disorders. Our findings emphasize PE as important risk markers for mental disorders during mid-adolescence and highlight the importance of monitoring children with PE before age 7 who develop persistent symptoms.
Langbaum, J. B.; Erickson, C. M.; Langlois, C.; Wood, E. M.; Egleston, B. L.; Harkins, K.; Mim, R.; John, S.; Brown, C.; Brown, S.; Howe, S.; Cacioppo, C.; Eppelmann, L.; Enos, J.; Salata, H.; DeSantiago, D.; Largent, E. A.; Reiman, E. M.; Denkinger, M. N.; Ashton, N. J.; Roberts, J. S.; Karlawish, J.; Bradbury, A. R.
Show abstract
Importance: Patients are increasingly learning Alzheimers disease (AD) genetic and biomarker results through electronic health portals. Evaluation of alternative scalable delivery models for return of AD risk information is needed to best support patient understanding and psychological well-being. Objective: To determine whether a patient-centered digital platform is comparable to clinician-mediated telehealth sessions for returning APOE and plasma pTau-217 results on outcomes of knowledge and psychological well-being. Design: The Evaluation of Self-Mediated Alternatives for Risk Testing Education and Return of Results (eSMARTER) study was a noninferiority trial of a patient-centered digital platform compared to clinician-mediated disclosure of APOE genotype and optional pTau-217 disclosure. Setting: Decentralized, fully remote trial enrolled participants in the contiguous United States (U.S.) between October 2024 and February 2025, with follow-up completed in November 2025. Participants: Eligible participants were aged 60-80 and had previously undergone APOE genotyping (without disclosure) via the GeneMatch program, passed psychological screening, had internet access, and were English-speaking. Interventions: Participants were randomized, 2:1, to the eSMARTER digital platform or clinician-mediated disclosure of APOE genotype. Following the 6-month post-APOE assessment, participants were offered optional pTau-217 disclosure via the same randomized modality. Main Outcomes and Measures: Primary outcomes at 1-7 days following APOE disclosure included changes in anxiety, disease-specific distress, and AD-related knowledge within a priori non-inferiority margins. Results: 674 persons (mean [SD] age 68 [4.7] years; 451 [67%] female; mean [SD] telephone MoCA=19 [2]) were eligible and provided demographic information. 651 participants were randomized to clinician-mediated (n=216) or digital disclosure (n=435) and completed APOE disclosure (66 [10%] APOE4 homozygotes, 377 [58%] heterozygotes, 208 [32%] non-carriers). 604 participants completed the study; 500 completed optional pTau-217 disclosure. Baseline characteristics were balanced across groups. At 1-7 days following APOE disclosure, scores on AD-related knowledge, PROMIS Anxiety, and disease-specific distress measures met non-inferiority. Conclusions and Relevance: Disclosure of APOE genotype by the eSMARTER digital platform is non-inferior to clinician-mediated telehealth disclosure. No significant between group differences were found following disclosure of pTau-217 results. Together, these results suggest that this digital platform may provide an evidence-based scalable approach for returning AD genetic and biomarker results.
Zink, T.; Noren, H.; Valdivia, D.; Yohn, C.; Hundal, J.; Chen, S.; Scarisbrick, D.; Sun, H.
Show abstract
Abstract: Objective: Post-traumatic epilepsy (PTE) is a common sequela of traumatic brain injury (TBI). Research indicates that individuals with PTE tend to experience greater cognitive difficulties compared to those with TBI alone. However, it is plausible that a distinct cognitive profile exists that distinguishes between TBI cases with and without PTE. We aimed to identify longitudinal changes in cognitive measures among TBI patients to better assess the changes associated with developing PTE. Setting: Outpatient. Participants: Prospective subjects who had suffered TBI within 6 months post-injury (TBI-6M, n=32), retrospective subjects with pre-existing PTE diagnoses (PTE, n=20), and healthy control subjects (HC, n=41). Design: We examined cognitive performance for TBI patients within 6 months post-injury, then again within 12 months (TBI-12M, n=26), and within 18-months (TBI-18M, n=25), and compared this with cognitive performance among HC and PTE. Main Measures: Cognitive tests administered yielded 15 test components for analysis. We utilized linear mixed effects modeling to examine cohort-level differences cognitive function. Results: 11/15 tests showed a significant performance deficit in the PTE subjects compared to HC. TBI-6M was not significantly different from the PTE subjects; with time, 9/15 tests showed some degree of recovery in TBI subjects. Tests for information processing speed/working memory and executive function showed strong recovery (TBI-6M vs. TBI-18M, SDMT written: p<0.0001, SDMT oral and COWAT: p<0.001). Tests for visual attention/working memory also showed a smaller but significant recovery (TBI-18M vs. PTE, p<0.05). By contrast, tests for verbal memory [HVLT-R Delayed Recall] showed chronic impairment in TBI (TBI-18M vs HC, p<0.0001). TBI subjects generally trend towards recovery in cognitive performance post-TBI. Conclusions: Information processing speed/working memory are strong indicators for TBI recovery, while auditory learning/memory shows chronic impairment. The stagnation of recovery in cognitive domains typically characterized by robust recovery may correlate with an elevated risk of developing PTE.
Ndiaye, A.; Thiebaut, A. C. M.; Borel, P.; Sabran, C.; Elis, S.; Guerif, F.; Maillard, V.
Show abstract
The distribution of fat-soluble compounds (including antioxidants) in follicular fluid (FF) remains sparsely documented in relation to in vitro fertilization (IVF) outcomes and existing studies have reported diverging associations. This study aimed to describe plasma and FF concentrations of fat-soluble micronutrients in women undergoing IVF and to analyze their adjusted associations with ovarian function, embryo development and pregnancy outcomes. In 2021-2022, plasma and FF samples were collected from 82 women (first IVF cycle) at oocyte puncture, along with lifestyle data covering the three preceding months. Eleven compounds (two tocopherols, three xanthophylls, five carotenes and retinol) were quantified. All compounds were detected in both compartments (lowest in FF) except phytoene, undetectable in FF. Plasma and FF -tocopherol concentrations were positively associated with plasma estradiol levels before oocyte puncture (both p<0.01) while FF -carotene and lycopene were inversely associated with plasma progesterone concentrations (p=0.01 and 0.02, respectively). Plasma phytofluene and phytoene were positively associated with mature oocyte rate (p=0.03 and p=0.01, respectively), while FF retinol was negatively associated (p=0.03). Carotenes, tocopherols and retinol were inversely associated with later IVF outcomes: fertilization rate (p<0.001 for plasma g-tocopherol, 0.02 for FF retinol), top-quality embryo (p=0.02 for plasma phytofluene), biochemical pregnancy at day 7 post-embryo transfer (p=0.05 for plasma -tocopherol, 0.02 for plasma -carotene), clinical pregnancy (p=0.03 for plasma -tocopherol, 0.01 for plasma phytoene) and live birth (p=0.04 for plasma -tocopherol, 0.02 for plasma phytoene). Plasma and FF g-tocopherol were positively associated with embryo fragmentation (both p<0.05). Finally, among xanthophylls, only plasma {beta}-cryptoxanthin was positively associated with plasma progesterone concentrations (p=0.02). Our findings of heterogeneous associations between tocopherols, carotenes, retinol and IVF outcomes across the stages of IVF suggest a beneficial effect limited to early outcomes and support a complex and context-dependent role of these compounds in female reproduction. This manuscript has been submitted to PlosOne on August 19, 2026.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.
Show abstract
Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.